Papers with lightweight training-based approach
You Only Need One Single Token to Refine Safety Alignment (2026.findings-acl)
Copied to clipboard
| Challenge: | Excessive safety can lead to over-refusal, where models reject harmful-looking yet benign queries, severely limiting utility. |
| Approach: | They propose a lightweight training-based approach that reshapes the distributions of harmful and benign samples within the model’s decision space by using a single-token prefix. |
| Outcome: | The proposed approach can distinguish between harmful and benign samples while keeping the model frozen. |